Papers with NLP technology

8 papers
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)

Copied to clipboard

Challenge: In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks.
Approach: They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia.
Outcome: The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons.
Cross-lingual Few-Shot Learning on Unseen Languages (2022.aacl-main)

Copied to clipboard

Challenge: Large pre-trained language models have demonstrated the ability to obtain good performance on downstream tasks with limited examples in resource-rich languages.
Approach: They propose to use a downstream sentiment analysis task to analyze the effectiveness of several few-shot learning strategies across 12 languages, including 8 unseen languages, to compare results.
Outcome: The proposed model, XLM-R, gives the best performance on a task with few examples in resource-rich languages.
Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond Gender (2022.coling-1)

Copied to clipboard

Challenge: Current modeling of 3rd person pronouns ignores neopronoun phenomena like naive pronounes, which are not (yet) widely established.
Approach: They propose to validate existing and novel approaches for modeling 3rd person pronouns in language technology and validate them through a survey.
Outcome: The proposed model excludes non-binary users, while ignoring gender-specific phenomena.
Evaluating the Diversity, Equity, and Inclusion of NLP Technology: A Case Study for Indian Languages (2023.findings-eacl)

Copied to clipboard

Challenge: In order for NLP technology to be widely applicable, fair, and useful, it needs to serve a diverse set of speakers across the world’s languages, be equitable, not unduly biased towards any particular language, and be inclusive of all users.
Approach: They propose to use Gini coefficient to assess NLP across all three dimensions to assess diversity, equity, and inclusion across all languages.
Outcome: The proposed evaluation paradigm assesses NLP technologies across all three dimensions and identifies the need for regional-specific choices in model building and dataset creation.
Benchmarking the Simplification of Dutch Municipal Text (2024.lrec-main)

Copied to clipboard

Challenge: Text simplification (TS) is a technique that makes written information more accessible to all people, especially those with cognitive or language impairments.
Approach: They propose to use English as a pivot language for simplification of Dutch medical and municipal texts.
Outcome: The proposed approach improves on Dutch medical text, while the existing pipeline performs better on all metrics.
NERetrieve: Dataset for Next Generation Named Entity Recognition and Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a widely adopted NLP task . authors present three variants of NER task, with dataset to support them .
Approach: They propose three variants of the NER task, together with a dataset to support them . they propose a move towards more fine-grained entities and zero-shot recognition .
Outcome: The proposed model matches or surpasses existing models in NER tasks . the proposed model is based on a large, silver-annotated corpus of 4 million paragraphs .
GlotLID: Language Identification for Low-Resource Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing web-mined datasets for low-resource languages have been useful for low resource NLP.
Approach: They propose a model that identifies 1665 low-resource languages and a new model that is rigorously evaluated and reliable.
Outcome: The proposed model outperforms baselines when balancing F1 and false positive rate (FPR).
Impoverished Language Technology: The Lack of (Social) Class in NLP (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on socio-demographic factors has focused on how much a person's socioeconomic status affects their language production and perception.
Approach: They propose to include socio-economic class in future natural language processing (NLP) research aimed at understanding relationships between socio-demographic factors and language production and perception.
Outcome: The proposed definition of class can be operationalised by NLP researchers and argue for including socio-economic class in future language technologies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations